Papers by Ngoc Thang Vu

27 papers
Exploring Segmentation Approaches for Neural Machine Translation of Code-Switched Egyptian Arabic-English Text (2023.eacl-main)

Copied to clipboard

Challenge: Code-switching (CS) is a problem in machine translation, but its performance is not investigated for CS settings.
Approach: They propose to use morphological segmentation techniques for machine translation tasks . they compare morphology-based and frequency-based segmentation for MT tasks based on data size .
Outcome: The proposed approach performs best in MT tasks but under-performs in other languages.
ADVISER: A Dialog System Framework for Education & Research (P19-3)

Copied to clipboard

Challenge: In this paper, we focus on task-oriented dialog systems, although our framework allows easy integration of non-task dialog systems and their combination.
Approach: They propose an open source dialog system framework for education and research that supports multi-domain task-oriented conversations in two languages.
Outcome: The proposed framework supports multi-domain task-oriented conversations in two languages and is open source for education and research.
Toward Implicit Reference in Dialog: A Survey of Methods and Data (2022.aacl-main)

Copied to clipboard

Challenge: In natural language, speakers often leave out information that is understood by the other party through the shared context.
Approach: They propose to use omitted entities as implicit references in dialogs to improve language processing.
Outcome: The proposed method is based on a set of experiments which show that the proposed method has a high level of accuracy and is a success.
F1 is Not Enough! Models and Evaluation Towards User-Centered Explainable Question Answering (2020.emnlp-main)

Copied to clipboard

Challenge: Existing models and evaluation settings have shortcomings regarding the coupling of answer and explanation which might cause serious issues in user experience.
Approach: They propose a hierarchical model and a new regularization term to strengthen the coupling of answer and explanation and two evaluation scores to quantify the couple.
Outcome: The proposed model strengthens the answer-explanation coupling and provides evaluation scores that align with user experience.
Beyond Accuracy: A Consolidated Tool for Visual Question Answering Benchmarking (2021.emnlp-demo)

Copied to clipboard

Challenge: Existing evaluation tools for general Visual Question Answering (VQA) systems are limited to answering accuracy, but they can be used to evaluate performance in real-world scenarios.
Approach: They propose a browser-based benchmarking tool with an API for easy integration of new models and datasets to keep up with the fast-changing landscape of VQA.
Outcome: The proposed tool tests generalization capabilities of models across multiple datasets and includes metrics that measure biases and uncertainty to further explain model behavior.
Neighboring Words Affect Human Interpretation of Saliency Explanations (2023.findings-acl)

Copied to clipboard

Challenge: Recent studies found that superficial factors such as word length can distort human interpretation of the communicated saliency scores.
Approach: They conduct a user study to examine how the marking of a word’s *neighboring words* affect the explainee’s perception of the word’ s importance in the context of . a saliency explanation.
Outcome: The findings question whether text-based saliency explanations should continue to be communicated at word level and inform future research on alternative methods.
Fine-tuning BERT for Low-Resource Natural Language Understanding via Active Learning (2020.coling-main)

Copied to clipboard

Challenge: Recent work has explored the suitability of pre-trained language models in low resource settings with less than 1,000 training data points.
Approach: They propose to use pool-based active learning to speed up training while keeping the cost of labeling new data constant.
Outcome: The proposed model can be fine-tuned to optimize for low-resource settings while keeping the cost of labeling constant.
»textklang« – Towards a Multi-Modal Exploration Platform for German Poetry (2022.lrec-1)

Copied to clipboard

Challenge: »textklang« aims to explore the relationship between written text and its potential and actual sonic realisation in lyric poetry . the platform will combine three modalities: the poetic text, the audio signal of a recorded recitation and, at a later stage, music scores of . musical setting of lyrical poetry.
Approach: They propose to combine a multi-modal corpus of German lyric poetry from the Romantic era with a platform for systematic exploration.
Outcome: The platform will combine the poetic text, the audio signal of a recorded recitation and, at a later stage, music scores of . a musical setting of lyric poetry.
Meta Learning and Its Applications to Natural Language Processing (2021.acl-tutorials)

Copied to clipboard

Challenge: Meta-learning is a new technique that aims to learn better learning algorithms, including better parameter initialization, optimization strategy, network architecture, distance metrics, and beyond.
Approach: This tutorial introduces Meta-learning approaches and the theory behind them, and then reviews the works of applying this technology to NLP problems.
Outcome: This tutorial will introduce Meta-learning approaches and the theory behind them, and then review the works of applying this technology to NLP problems.
Prompting-based Synthetic Data Generation for Few-Shot Question Answering (2024.lrec-main)

Copied to clipboard

Challenge: Language models have boosted the performance of Question Answering, but data annotation is costly.
Approach: They propose to use large language models to improve Question Answering performance . they argue that domain-agnostic knowledge from LMs is sufficient to create a well-curated dataset.
Outcome: The proposed model outperforms state-of-the-art approaches on few-shot Question Answering.
A Survey of Code-switched Arabic NLP: Progress, Challenges, and Future Directions (2025.coling-main)

Copied to clipboard

Challenge: Code-switching (CSW) is a common linguistic phenomenon in multilingual societies . current literature on CSW in the arab world is limited to the Arabic language .
Approach: They present a review of the literature in the field of code-switched Arabic NLP . they propose recommendations for future research .
Outcome: This review provides a broad perspective on the current literature in the field of code-switched Arabic NLP . it also provides recommendations for future research .
Intrinsic Subgraph Generation for Interpretable Graph Based Visual Question Answering (2024.lrec-main)

Copied to clipboard

Challenge: Visual Question Answering (VQA) is acknowledged as a challenging multi-modal task for Machine Learning (ML).
Approach: They propose an interpretable approach for graph-based Visual Question Answering . their model is designed to intrinsically produce a subgraph during the question-answering process as its explanation .
Outcome: The proposed model outperforms existing explainable methods on a graph-based VQA dataset.
Ethical Considerations for Machine Translation of Indigenous Languages: Giving a Voice to the Speakers (2023.acl-long)

Copied to clipboard

Challenge: In recent years, machine translation has become very successful for high-resource language pairs.
Approach: They conduct interviews with community leaders, teachers, and language activists to shed light on ethical considerations for the automatic translation of Indigenous languages.
Outcome: The results show that the inclusion of native speakers and community members is vital to performing better and more ethical research on Indigenous languages.
Introducing Two Vietnamese Datasets for Evaluating Semantic Models of (Dis-)Similarity and Relatedness (N18-2)

Copied to clipboard

Challenge: Existing datasets for low-resource language Vietnamese assess semantic similarity . a dataset for word pairs with similarity levels is needed to evaluate these models .
Approach: They present two new datasets for the low-resource language Vietnamese to assess models of semantic similarity.
Outcome: The two datasets are comparable to the English datasets.
Low-Resource Multilingual and Zero-Shot Multispeaker TTS (2022.aacl-main)

Copied to clipboard

Challenge: Currently, the amount of data needed for TTS is limited to the vast majority of the spoken languages.
Approach: They propose to use language agnostic meta learning procedure to learn speaking a new language with just 5 minutes of training data while retaining the ability to infer the voice of even unseen speakers.
Outcome: The proposed approach is able to learn speaking a new language using just 5 minutes of training data while retaining the ability to infer the voice of even unseen speakers in the newly learned language.
Cairo Student Code-Switch (CSCS) Corpus: An Annotated Egyptian Arabic-English Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Code-switching is a phenomenon commonly observed in the Arabicspeaking world . there is still a huge gap in the available resources and NLP applications .
Approach: They propose a corpus of Egyptian- Arabic code-switch speech data that is fully tokenized, lemmatized and annotated for part-of-speech tags.
Outcome: The proposed corpus of Egyptian- Arabic code-switch speech data is fully tokenized, lemmatized and annotated for part-of-speech tags.
Fast and Accurate Non-Projective Dependency Tree Linearization (2020.acl-main)

Copied to clipboard

Challenge: Existing methods for decoding dependency trees are 10 times faster than current ones.
Approach: They propose a graph-based method to tackle a dependency tree linearization task . they propose to solve a Traveling Salesman Problem and combine the solution into a projective tree .
Outcome: The proposed method outperforms the state-of-the-art linearizer while being 10 times faster in training and decoding.
Few-shot Learning for Slot Tagging with Attentive Relational Network (2021.eacl-main)

Copied to clipboard

Challenge: Recent studies have used metric-based learning in computer vision but not slot tagging.
Approach: They propose a metric-based learning architecture that extends relation networks by leveraging pretrained contextual embeddings such as ELMO and BERT and by using attention mechanism.
Outcome: The proposed method outperforms state-of-the-art methods on SNIPS data on a slot tagging task with a large amount of hand-labeled data.
ADVISER: A Toolkit for Developing Multi-modal, Multi-domain and Socially-engaged Conversational Agents (2020.acl-demos)

Copied to clipboard

Challenge: Existing toolkits for developing dialog systems are limited to core components and do not support multi-modal processing and social signals.
Approach: They propose to use ADVISER to develop multi-modal dialog agents using multi-text and social signals.
Outcome: The proposed toolkit is flexible, easy to use, and easy to extend for linguists and cognitive scientists, thereby providing a flexible platform for collaborative research.
Conversational Tree Search: A New Hybrid Dialog Task (2023.eacl-main)

Copied to clipboard

Challenge: Existing conversational interfaces are limited to FAQs and dialogs, allowing users to search for specific questions.
Approach: They propose a task that bridges the gap between FAQ-style information retrieval and task-oriented dialog.
Outcome: The proposed task bridges the gap between FAQ-style information retrieval and task-oriented dialog.
Explaining Pre-Trained Language Models with Attribution Scores: An Analysis in Low-Resource Settings (2024.lrec-main)

Copied to clipboard

Challenge: Currently, prompt-based models are gaining popularity due to their easier adaptability in low-resource settings.
Approach: They analyze attribution scores extracted from prompt-based models w.r.t. plausibility and faithfulness and compare them with attribution score extracted from fine-tuned models and large language models.
Outcome: The proposed model outperforms attention and Integrated Gradients in plausibility and faithfulness, while fine-tuning models are harder to explain in low-resource settings.
ArzEn: A Speech Corpus for Code-switched Egyptian Arabic-English (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of Arabic-English code-switching (CS) spontaneous speech is collected in an Egyptian university soundproof room . the language in Egypt is rather complex and poses many challenges to natural language processing (NLP)
Approach: They present an Egyptian Arabic-English code-switching (CS) spontaneous speech corpus.
Outcome: The proposed corpus is designed to be used in automatic speech recognition systems . it provides a useful resource for analyzing the CS phenomenon from linguistic, sociological, and psychological perspectives.
IMSurReal: IMS at the Surface Realization Shared Task 2019 (D19-63)

Copied to clipboard

Challenge: a system for shallow and deep completion is presented for the Surface Realization Shared Task 2019 . the system achieves state-of-the-art performance without using external data.
Approach: They propose a surface realization system that takes five steps without external data . they perform detailed error analysis revealing correlation between word order freedom and difficulty .
Outcome: The proposed system achieves state-of-the-art without external data . it achieves highest BLEU scores on tokenized text and human evaluation on four languages .
Discrete Subgraph Sampling for Interpretable Graph based Visual Question Answering (2025.coling-main)

Copied to clipboard

Challenge: XAI aims to make machine learning models more transparent, but interpretable approaches are relatively rare.
Approach: They integrate discrete subset sampling methods into a graph-based visual question answering system to evaluate their interpretability.
Outcome: The proposed methods mitigate trade-off between interpretability and answer accuracy while achieving strong co-occurrences between answer and question tokens.
It’s What You Say and How You Say It: Investigating the Effect of Linguistic vs. Behavioral Adaptation in Task-Oriented Chatbots (2025.coling-main)

Copied to clipboard

Challenge: linguistic adaptation is not known to have a positive impact on dialog success and user perception.
Approach: They evaluate subjective and objective aspects of dialog success and user perceptions through a user study . they also examine linguistic adaptations of dialog agents to determine which aspects influence user perception .
Outcome: The proposed agents can differ in their level of formality and their linguistic style.
DIAGRAPH: An Open-Source Graphic Interface for Dialog Flow Design (2023.acl-demo)

Copied to clipboard

Challenge: Dialog systems have gained attention as a convenient way for users to access information in a more personalized manner.
Approach: They present a graphical dialog flow editor built on ADVISER toolkit . it provides a clean and intuitive graphical interface for creating dialog systems .
Outcome: The tool is based on the ADVISER toolkit and is evaluated with subject-experts . it is able to quickly prototype dialog systems and provide a test bed for students learning about dialog systems.
Towards a Zero-Data, Controllable, Adaptive Dialog System (2024.lrec-main)

Copied to clipboard

Challenge: Recent approaches to controllable dialog systems require additional training data to be deployed in new domains.
Approach: They propose to generate dialog tree data directly from dialog trees by using a commercial Large Language Model or a single GPU.
Outcome: The proposed approach can achieve comparable dialog success to models trained on human data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations